Skip to content

[Fix] Fit the browser front's state and questions into Laya's window - #9

Open
chaimaerachdi wants to merge 3 commits into
ThinkFlowLab:mainfrom
chaimaerachdi:laya-compact-browser-state
Open

chaimaerachdi wants to merge 3 commits into
ThinkFlowLab:mainfrom
chaimaerachdi:laya-compact-browser-state

Conversation

@chaimaerachdi

@chaimaerachdi chaimaerachdi commented Sep 24, 2026 •

Copy link
Copy Markdown

Why

--model laya on a browser agent could not work as the front sent it. Two things did not fit Laya's window:

  1. The state. The browser front sends every model the state it sends Jev: a JSON object per element row, the
    full page text and ten actions of history, about 7K tokens per decision on Google Flights and 18,785 to 23,654
    on Allrecipes (docs/benchmarks.md). Laya reads 512 to 1024 tokens.
  2. The questions. Laya fits a question's instruction and all its options into one head_max_len budget (192
    by default). Past it, every option is cut to an equal share. The browser front's target options are JSON
    objects, so a 23-element click head left each option six tokens, 12: {"element": "[: Laya never saw an
    element's name. The instruction also carried the agent's rules, 446 tokens.

How

Both folds are in s1a/decision_models/laya.py and run in LayaModel._decide only when the state has the browser
front's shape; the tool front and the rails pass through unchanged.

  • laya_state() folds the state: page.text dropped, one short line per element row, the last three actions.
  • laya_browser_question() folds each question: the instruction becomes the goal and the operation, and each target
    option becomes its element's label and value, Where from? = Zurich.

What

  • LAYA_COMPACT_BROWSER_STATE (default on; 0, false or no turns both folds off).
  • Browser runs need a wider window than the checkpoint's default. Measured on Google Flights with this shape: 157 to
    206 tokens for the operation head, 73 to 101 for TYPE_TEXT and PRESS_ENTER, 92 to 915 for CLICK (a calendar page
    offers 66 days), 98 to 1,002 tokens of state. So: LAYA_MAX_LEN=1536 LAYA_HEAD_MAX_LEN=1024.
  • Docs: docs/decision-models.md, docs/configuration.md, .env.example, CHANGELOG.md.
  • [Feat] Integrate and evaluate Julia-1 #25 (Julia-1) plans to move laya_state() to shared code for Julia-1 as well.

Results: the request fits, the stock checkpoint still fails

Google Flights, Zurich to London, one adult, economy. Same Chrome, CPU only (no GPU).

model result decisions time per decision
Jev (reference, 2026-09-25) DONE, flights shown (€49 easyJet) 12 267 ms median
Laya, this PR, 1536/1024 (2026-09-28) DONE at the first step, nothing filled 1 12.4 s
Laya, state fold only (2026-09-25) DONE at step 0, or clicks "Munich" 1 to few 7 to 11 s
Cua-S1 Nano (2026-09-25) types into "Return", then stops few 121 to 132 ms
  • On the 2026-09-28 run Laya put DONE at 0.55 and CLICK at 0.23, over 1,254 input tokens for four questions.
  • The state-fold-only runs failed the same way with the fold off, and with the typed-decisions checkpoint.
  • Offline, replaying Jev's twelve recorded steps of the same task through this PR's shape, Laya picks Jev's answer
    on 4 of 23 questions (the operation head and the chosen operation's target head).

So the fit problem is fixed, but Laya does not know how to drive a web form. Its model card names email triage,
routing, guardrails and moderation as what its checkpoints are for. Browser use would need a checkpoint fine-tuned on
browser steps. I recorded Jev on 12 more Google Flights routes as training data (all DONE, 145 decisions), but
fine-tuning on my CPU was too slow to finish (about 30 minutes per pass over 278 examples). A GPU would make it
possible.

Verification

  • uv run ruff format --check . && uv run ruff check .: 137 files formatted, all checks passed.
  • uv run ty check: 1 diagnostic, in s1a/decision_models/cua.py, not touched here; the same on main.
  • uv run pytest -q --ignore=tests/test_browser_policy.py: 16 failed, 404 passed, 41 skipped. main on this
    machine: 16 failed, 390 passed. The same 12 failing tests by name on both (Windows environment), none new.
    tests/test_browser_policy.py fails to collect on main too (openjiuwen.harness.schema.decision_policy
    missing in the installed openjiuwen).
  • uv run pytest tests/test_decision_models_laya.py -q: 43 passed, 2 skipped (fake agent, no weights).
  • scripts/smoke.sh: smoke: ok.
  • CHANGELOG.md and the docs say what the code does now.
  • Real run: LAYA_MAX_LEN=1536 LAYA_HEAD_MAX_LEN=1024 uv run s1a run flights --model laya --timeout 900 (with
    uv sync --extra laya), result above.

🤖 Generated with Claude Code

chaimaerachdi and others added 2 commits September 28, 2026 12:30
Laya reads a 512 to 1024 token window, but the browser front sends
every model the same state it sends Jev: a JSON object per element
row, the full page text, and ten actions of history. On a real page
that state fills the window well before a single instruction token
is spent (docs/benchmarks.md shows 18,785-23,654 input tokens for
Jev on the Allrecipes run), which is why `--model laya` routinely
raises MODEL_SERVICE_CONFIG_ERROR on the browser front today.

laya_state() folds a browser-shaped state before every call to
LayaModel._decide: page.text dropped (the choice heads already carry
each candidate's own text; the free-form dump is for the chat
model's DONE answer, which Laya never writes), each element row
rendered as one short line instead of a JSON object, and the last
three actions kept instead of ten. On by default; anything that
isn't the browser front's shape passes through unchanged (the tool
front already fits). LAYA_COMPACT_BROWSER_STATE=0 turns it off.

34 unit tests (tests/test_decision_models_laya.py) cover the
compaction itself, its wiring into LayaModel, and the env-var
opt-out, all against FakeLayaAgent (no torch/weights needed). Full
suite run against main: identical 19 pre-existing failures before
and after this change (missing `ty` binary and other env-only gaps
in this sandbox, unrelated to decision_models/laya.py).

Not done here, and worth flagging: this closes the "state is bigger
than the window" gap, not the "is Laya's window, even filled,
actually fast enough end to end on a real page" question. That
needs the real convaiinnovations/laya checkpoint (uv sync --extra
laya), a live browser run, and a real latency number. See the PR
description for exact commands.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Laya fits a question's instruction and all its options into one head_max_len
budget (192 by default). The browser front's target options are JSON objects,
so a 23-element click head left each option about six tokens,
`12: {"element": "[`, and Laya never saw an element's name.

laya_browser_question rewrites each browser question when the state is folded:
the instruction becomes the goal and the operation (the agent's rules, 446
tokens, are dropped), and each target option becomes its element's label and
value. Docs and .env.example give the window a browser run needs
(LAYA_MAX_LEN=1536, LAYA_HEAD_MAX_LEN=1024) and the measured result: the stock
checkpoint still answers DONE at the first step on Google Flights.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
@chaimaerachdi
chaimaerachdi force-pushed the laya-compact-browser-state branch from 0c470d9 to e513234 Compare September 28, 2026 10:34
@chaimaerachdi chaimaerachdi changed the title [decision_models] Compact the browser front's state for Laya's window [Fix] Fit the browser front's state and questions into Laya's window Sep 28, 2026
@chaimaerachdi
chaimaerachdi marked this pull request as ready for review September 28, 2026 10:42
@sss-hust

Copy link
Copy Markdown

Thanks for this. The diagnosis is solid: the six-token options (12: {"element": "[) are a real problem, and it's good that the docs say plainly that the stock checkpoint still answers DONE at step 1. That keeps anyone from reading this as "Laya now drives the browser". I found one bug and two smaller issues about lost information.

1. text_value gets asked "Which operation comes next?" (bug)

laya_browser_question rewrites any choice question that has a goal. But text_value (s1a/browser/action_space.py:200-205) has a goal and no operation, so it falls through to the operation-head wording. Reproduced on e513234 with a small Google Flights-shaped snapshot through build_action_space / build_questions:

== text_value | Task: Find flights from Zurich to London Which operation comes next?
   {'London': 'London', 'Zurich': 'Zurich', 'none': 'No offered value fits the chosen field.'}

Laya is asked to choose an operation from a list of values. _decide has the question name, so one fix is to key the rewrite on the head (operation and *_target) and give text_value its own ask, e.g. "Which value should be typed into the field?". A test next to test_the_operation_head_keeps_its_string_options would pin it.

2. Truncation collapses distinct options into the same text

With 28 characters per option and the [index] prefix stripped, different targets become identical strings. The same snapshot gives:

click_target: {'2': 'Select flight: Swiss LX 318,', '3': 'Select flight: Swiss LX 318,', '4': 'Search', '5': 'Search', ...}
state rows:   "2 button Select flight: Swiss LX 318, departs 07:", "3 button Select flight: Swiss LX 318, departs 07:"

Result lists and calendars (where labels share a long prefix) are exactly the pages the docs measure. Does Laya read the criteria keys, or only the option text? If only the text, it can't tell these apart. Possible fixes: keep the tail of the label instead of the head when labels collide, or add the index or region only when two options would otherwise be equal.

3. blocked_by and region are dropped (minor)

_laya_browser_row reduces blocked_by to the letter B, and drops region from both rows and options. The browser rules (s1a/browser/prompts.py:30) tell the model to click the overlay's own button, and the overlay's name is the only link between the blocked control and that button. A truncated blocked_by name costs a few tokens. I don't think this blocks the PR.

Non-blocking

  • The "roughly tenfold" figure: can the measurement be reproduced from the repo (a script or the recorded pages)? If not, "on the pages measured" is fine, but a pointer would help the next person.
  • if state is not observation.state uses object identity to detect a browser state. It works, but a small is_browser_state(state) helper shared with laya_state would make the intent explicit.

What I ran (macOS 26.5, Python 3.14; CI uses 3.11/3.13): ruff format --check, ruff check, ty check all clean; pytest -q 505 passed, 41 skipped; scripts/smoke.sh ok. I did not run Laya itself (no laya extra). The repros above use the real build_action_space / build_questions with a hand-written snapshot.

Requesting changes for item 1; items 2 and 3 are up to you.

- text_value (a goal, no operation) was asked "Which operation comes next?".
  The rewrite now keys on the head: operation, <op>_target, text_value
  ("Which value should be typed into the field?").
- Options cut to 28 characters could become identical (a result list, a
  calendar); two options with the same text now keep their key in front.
- A blocked row kept only the letter B; it now keeps its overlay's name, the
  link to the button that closes it.
- is_browser_state() names the check laya_state and _decide share, in place
  of the object-identity test.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
@chaimaerachdi

Copy link
Copy Markdown
Author

Thanks for the careful review. Fixed in ecefeea:

  1. text_value now has its own ask ("Which value should be typed into the field?"); the rewrite is keyed on the head name.
  2. Options that shorten to the same text now keep their key in front ([2] ..., [3] ...).
  3. A blocked row keeps its overlay's name ("blocked by ...").
    Also added is_browser_state() instead of the identity check. The "tenfold" figure comes from the fixture in test_compaction_shrinks_the_json_size_by_an_order_of_magnitude. Tests added for each point

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants